Papers with Multimodal Chain of Thought
BQA: Body Language Question Answering Dataset for Video Large Language Models (2025.acl-short)
Copied to clipboard
| Challenge: | a large part of human communication relies on nonverbal cues such as facial expressions, eye contact, and body language. |
| Approach: | They propose to validate whether video large language models can correctly interpret body language from short clips of body language. |
| Outcome: | The proposed model can correctly interpret emotions from short clips of body language. |